Nature Machine Intelligence
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match Nature Machine Intelligence's content profile, based on 70 papers previously published here. The average preprint has a 0.10% match score for this journal, so anything above that is already an above-average fit.
Midjani, F.; Ghelich, R.; Keshtkar, F. Z.; Malekpour, M.; Lee, H.
Show abstract
Peptides are increasingly explored as therapeutic candidates, delivery vectors, and functional biomolecules, but experimental screening of peptide activity and safety remains costly because the sequence space is vast and small sequence changes can alter functionality. Computational peptide classification can therefore help prioritize candidates. However, many protein-language-model-based classifiers achieve strong performance using opaque prediction heads, making it difficult to determine which learned evidence supports or opposes a prediction. We present RulePep, an ESM-2-guided neural-symbolic classifier for peptide-function prediction. RulePep maps frozen ESM-2 sequence representation to learned latent predicates, polarity-constrained differentiable rules, and an additive symbolic logit whose components can be inspected at the case level. We evaluate RulePep on three biologically distinct peptide classification tasks: blood-brain barrier penetration, hemolytic potency, and anticancer activity. On the BBPpredict, HemoPI3, and AntiCP 2.0 alternate benchmark datasets, RulePep achieved AUROC/MCC values of 0.8869/0.6850, 0.9155/0.6820, and 0.9765/0.8633, respectively. Ablation experiments supported the contributions of multi-layer representation pooling, rule polarity, mined-rule initialization, symbolic capacity, and rule-derived aggregation. RulePep combines competitive predictive performance with additive logit reconstruction, rule-level evidence reporting, and predicate-suppression auditing, providing a transparent sequence-based framework for peptide candidate prioritization.
Steenwyk, J. L.
Show abstract
Protein language models learn general-purpose representations from large collections of protein sequences and structures, and have advanced the prediction of protein structure and function. ESM3 is a multimodal protein language model that ingests a protein through several channels at once, including amino-acid sequence, three-dimensional structure, secondary structure (SS8), solvent accessibility (SASA), and discrete functional annotations, summing their embeddings into a single residual stream. Little is known about whether these modalities occupy separate subspaces and the depth at which they fuse. The present analysis examines ESM3 (esm3-sm-open-v1; 1.4 billion parameters; 48 transformer layers) once per modality in isolation and applies representational-similarity analysis across all 48 layers. The four physical modalities (sequence, structure, SS8, SASA) begin in distinct subspaces, remain maximally separated through roughly the first half of layers, and then fuse into a shared low-dimensional subspace between layers 25 and 35. The fusion is ordered. The structure-derived modalities (structure, SS8, SASA) are mutually aligned from the input, whereas sequence joins last, after layer 28. The functional-annotation modality never fuses; instead, it remains representationally orthogonal to the physical modalities at every layer, and this orthogonality holds whether the annotation is supplied as whole-protein or per-residue, suggesting that it is content-driven rather than a tokenization arti-fact. The fusion is a learned property, absent in a randomly initialized model of the same architecture, holds at the residue level below the mean-pool, and reorganizes variance, converting between-condition variance into within-condition variance while the stream never approaches isotropy. Fusion depth is independent of protein length but is delayed by structural disorder. The phenomenon is universal across diverse organisms. Across 5,555 proteins from 12 organisms spanning eukaryota, bacteria, and archaea, every superkingdom (and every individual organism) reaches peak modality fusion at the same network depth (layer 35).
Long, W.; Liu, T.; Szalata, A.; Theis, F. J.; Xue, L.; Zhao, H.
Show abstract
Multimodal single cell perturbation screens offer a scalable approach for characterizing the effects of genetic and chemical interventions on cellular state. However, most existing representation learning methods are tailored to a single perturbation modality and fail to explicitly incorporate external semantic knowledge, which limits their ability to generalize across datasets and perturbation types. Here, we introduce PertOmni, a CLIP style multimodal representation learning framework that aligns transcriptomic perturbation signatures with text derived embeddings of curated genes and compound descriptions, as well as image derived embeddings from cell paintings. PertOmni jointly trains a shared transcriptomic encoder and dataset specific text encoders using a masked contrastive objective that emphasizes within cell type discrimination while mitigating confounding effects arising from cell type heterogeneity. We evaluate the produced joint embedding space on bidirectional retrieval, drug gene interaction inference, and perturbation prediction across both small molecule and CRISPRi perturbation datasets, and demonstrate consistent improvements over strong baseline methods.
Han, R.; Yu, B.; Xinghui, S.; Xiao, L.; Junhai, Q.; Ting, Y.; Xin, G.
Show abstract
Nanopore direct RNA sequencing enables direct profiling of RNA modifications on native transcripts, but accurate multi-modification detection remains limited by non-stationary signals and heterogeneity across chemistries. Here, we develop WattmaMod, a deep learning framework for multi-modification detection from nanopore direct RNA sequencing data. It combines self-supervised pretraining, supervised contrastive fine-tuning, and low-label incremental adaptation to improve representation learning and support efficient extension to low-resource modification types. The framework further incorporates wavelet-guided multi-scale encoding and dynamic cross-attention fusion to model raw signals and event-level features. Results show that WattmaMod achieves robust detection of multiple RNA modifications, including m6A, m5C, m1A, A-to-I, m7G, hm5C, m1{Psi}, f5C, ac4C, m5U and {Psi}. It also extends efficiently to low-resource modification types with minimal labeled data, generalizes across sequencing chemistries and species, and predicts potential higher-order local organization among distinct RNA modifications. WattmaMod thus provides a scalable framework for high-resolution epitranscriptome profiling and expands RNA modification analysis beyond single-site prediction to coordinated multi-modification characterization.
Zhang, C.; Li, H.; Tian, F.; Mansour L., S.; Orban, C.; Chen, C.; Zhou, J. H.; Yeo, B. T. T.; the Alzheimer's Disease Neuroimaging Initiative, ; the Australian Imaging Biomarkers and Lifestyle Study of Ageing,
Show abstract
Longitudinal dementia progression prediction is essential for clinical decision-making. However, models often degrade on external cohorts due to systemic missingness -- where certain biomarkers available during training are completely absent at test time -- compounded by distribution shifts and patient-specific variability. Here, we propose Progression-aware Feature Fusion with Test-Time Adaptation (ProFuse-TTA), a two-stage hierarchical Transformer for longitudinal dementia prediction. Stage 1 learns per-biomarker temporal representations from irregular observations without imputation. Stage 2 fuses them via cross-feature attention, with simulated modality dropout during training for robustness to systemic missingness. At inference, a lightweight test-time adaptation module performs per-individual calibration. We trained on ADNI and evaluated on three external cohorts comprising 2,316 participants and 13,205 timepoints, with controlled modality ablation experiments isolating the effect of systemic missingness. We compared against six baselines, four from a recent benchmark study and two new baselines including one built on a tabular foundation model. ProFuse-TTA achieved the best cross-dataset performance in 8 of 9 settings across clinical diagnosis, MMSE, and hippocampal volume prediction, and ranked first in 14 of 15 ablation scenarios. The model maintained superior performance across varying input lengths and prediction horizons up to 6 years. Pretrained ADNI models are available at XXX.
Mahtabi, B.; Nasr-Esfahani, E.; Yaraghi, S.
Show abstract
Pneumonia is a leading cause of infectious disease mortality worldwide, accounting for approximately 2.5 million deaths annually and 15% of deaths in children under five. Chest X-ray imaging remains the primary diagnostic tool, but accurate interpretation requires radiological expertise that is disproportionately concentrated in high-income settings, creating a diagnostic gap where disease burden is highest. Automated deep learning offers a scalable complement to specialist-dependent diagnosis, yet clinical adoption requires both high accuracy and transparent, interpretable reasoning. Convolutional neural networks (CNNs) have shown strong potential for pneumonia detection from chest X-rays, but two barriers impede clinical translation: the interpretability of black-box models and the computational feasibility of large architectures in resource-constrained settings. Explainable AI (XAI) methods such as Grad-CAM, Grad-CAM++, and Score-CAM address the interpretability barrier, yet systematic quantitative comparisons across multiple CNN architectures remain scarce. Furthermore, CNN architectures widely used for medical image classification carry high parameter counts that limit feasibility in resource-constrained settings, motivating architectures that achieve competitive accuracy with substantially fewer parameters. Here we propose a parameter-efficient deep learning framework for pneumonia detection based on transfer learning, evaluated across three CNN architectures representing distinct architectural families: EfficientNet-B0 with fine-tuning (proposed method), ResNet50, and DenseNet121, trained under identical conditions on the Kaggle chest X-ray dataset (5,863 images). Our method achieved 90% classification accuracy, outperforming both baselines while requiring 4.8x fewer parameters than ResNet50. To evaluate explainability, Grad-CAM, Grad-CAM++, and Score-CAM were applied across all three architectures and compared quantitatively using Intersection over Union against manually annotated lung segmentation masks, Insertion score, and Deletion score, with pairwise statistical validation via Wilcoxon signed-rank tests and Bonferroni correction. Findings show that classification accuracy and XAI explanation quality must be evaluated independently, and that the proposed parameter-efficient architecture offers a favorable trade-off for resource-constrained clinical deployment.
Jia, S.; Lysenko, A.; Boroevich, K. A.; Sharma, A.; Tsunoda, T.
Show abstract
DeepInsight-style methods make tabular feature relationships accessible to convolutional networks by placing each feature at a fixed position on an image carrier. An open design question is how the carrier geometry should be constructed when feature neighborhoods themselves carry part of the signal. We introduce vDeepInsight, an injective three-dimensional (3D) voxel carrier that preserves feature neighborhoods more faithfully than matched two-dimensional (2D) carriers while keeping a one-to-one mapping from each feature to a single voxel. Genes are embedded with t-SNE or UMAP, assigned one-to-one to a sparse voxel grid by linear-sum assignment, and processed by a submanifold sparse 3D convolutional network. We evaluate the carrier on gene expression through four linked analyses. First, representation-quality metrics show that 3D layouts reduce gene-neighborhood distortion relative to matched 2D layouts before any model is trained. Second, controlled synthetic tasks show that a sparse 3D convolution can exploit this preserved locality, but only when the supervised signal is constructed to depend on co-located genes and the receptive field spans adjacent voxels. Third, on real omics tasks the 3D carrier matches or exceeds tuned tabular baselines and consistently exceeds matched 2D carriers; the margin is small on marker-type classification, where individual genes already carry much of the label (tissue, lineage and cancer-type classification), and larger on program-type tasks, where the target depends on coordinated, pathway-level multi-gene activity (drug-response regression, TCGA immunogenomic-context regression and mechanism-of-action classification). Fourth, because the assignment is injective, voxel attribution maps directly back to genes, enabling gene-level attribution and pathway-level functional interpretation without voxel-to-gene deconvolution. Overall, the added carrier dimension improves the fidelity of feature-neighborhood representation and translates this improvement into prediction gains that are largest when the signal is distributed across local gene programs rather than dominated by individual marker genes.
Azbijari, N.; Wynne, J. H.; David, M.; Thurber, A. R.
Show abstract
Since the early adoption of metagenomics (the culture-free sequencing of microbial community genomes) in 2011, sequence data has increased over 500-fold across ecosystems. This surge in data has outpaced reliable taxonomic and functional annotation, with over half of sequences lacking confident functional assignment. These unknown sequences limit our understanding of microbial processes central to planetary health and human health. Recent advances in genomic language modeling have made progress in the interpretation of metagenomics datasets. Most state-of-the-art models rely on transformer architectures, which limit the maximum sequence length and therefore capture only a fraction of assembled metagenomic sequences due to the quadratic scaling of attention. This prevents training and inference on sequences with broad context, including multiple coding and non-coding regions. To overcome this limitation, we propose leveraging new model architectures that scale linearly with sequence length, making them more suitable for modeling longer metagenomic sequences. Here, we introduce Nammu, a mixed-modality Mamba-based foundation model with 167M parameters trained on the OpenMetaGenomic (OMG) corpus. Nammu is a bidirectional encoder trained with a 20K context length using a curriculum strategy, first on 64M protein sequences and then on 32M mixed-modality metagenomic contigs. We compared Nammu to gLM2, a mixed-modality transformer also trained on OMG using 37% more tokens, using taxonomy inference on a marine dataset from the Critical Assessment of Metagenome Interpretation (CAMI). Nammu outperforms gLM2 at every taxonomic level. We further assessed function via KEGG Orthology prediction in deep-sea metagenome-assembled genomes, where Nammu outperforms gLM2 (150M). These results demonstrate improved performance.
Zhou, Q.; Le, Y.; Qi, X.; Chang, S.; Lu, H.; Wu, Y.; Wang, H.; Ran, R.; li, x.
Show abstract
Foundation models learned from single-cell transcriptomes are central to the prospect of AI virtual cell that can represent, query and predict cellular state. However, most current single-cell foundation models learn from a single view of gene expression and are optimized primarily through reconstruction or next-token prediction. As a result, they capture expression abundance but cannot explicitly reconcile complementary views of cellular state. Here we present CellOS, a multi-view foundation model that learns cellular representations from paired expression and perception views. CellOS integrates complementary views through a scalable three-stage training strategy that combines causal cell-sentence language modelling, function-preserving dense-to-mixture-of-experts expansion and latent-space alignment via an LLM-JEPA objective. Using this framework, we trained a 12-billion-parameter model on 390.5 million single-cell transcriptomes. Across diverse benchmarks spanning cell-state annotation, batch integration and perturbation-response prediction, CellOS consistently outperformed state-of-the-art single-cell foundation models. Together, these results suggest that predictive alignment between complementary cellular views provides a scalable path toward representation-centric cellular world models and transferable AI virtual cells.
Wang, R.; Jin, K.; Pan, L.
Show abstract
We investigate whether representations from AINN-P1--a protein foundation model trained autoregressively on tens of millions of natural protein sequences--transfer to the task of ranking antibody-antigen pairs by binding affinity. Casting affinity maturation as a learning-to-rank problem over the change in binding free energy ({Delta}{Delta}G), we compare a task-specific sequence model trained end-to-end from scratch against lightweight downstream heads built on top of frozen AINN-P1 embeddings, all evaluated under an identical five-fold cross-validation protocol. A regularized linear probe on the frozen embeddings already surpasses the from-scratch baseline, and an optimized lightweight head raises the mean Spearman rank correlation from 0.42 to 0.53--a relative improvement of approximately 28%-- while training in seconds and without any fine-tuning of the foundation model. Because a linear probe alone exceeds the fully trained end-to-end baseline, the gain is attributable to representation quality rather than to added downstream-model capacity. These results position frozen foundation-model embeddings as a strong, data-efficient default for affinity ranking in antibody engineering and establish a conservative lower bound that task-adaptive fine-tuning is expected to exceed.
Shenoy, A. R.; Mendez, T.
Show abstract
Stroke is a leading cause of death and long-term disability worldwide, affecting approximately 15 million individuals annually. Prompt and accurate subtype differentiation between ischemic and hemorrhagic stroke is clinically critical, as the two conditions demand diametrically opposite interventions - thrombolytic therapy versus surgical decompression. Yet the majority of existing deep learning approaches reduce this problem to binary detection, and virtually none address the opacity of their decision-making in a clinically actionable manner. We present CerebAI, an explainable, deployment-oriented three-class CT stroke classification system built on a fine-tuned ConvNeXt-Base backbone with Integrated Gradients (IG) attribution. Trained on 6,774 non-contrast CT scans stratified across No Stroke, Ischemic Stroke, and Hemorrhagic Stroke, CerebAI achieves a weighted F1-score of 0.9746 (95% CI: [0.9625, 0.9851]), accuracy of 97.47%, macro-averaged AUC of 0.9921, mean Intersection-over-Union (mIoU) of 0.9276, Expected Calibration Error (ECE) of 0.0115, mean Brier Score of 0.0150, and Cohen's {kappa} of 0.9483 - surpassing ResNet-50, EfficientNet-B4, and Vision Transformer (ViT-B/16) baselines across all reported metrics. Integrated Gradients produce pixel-precise saliency maps that localize pathological regions with greater anatomical fidelity than Gradient-weighted Class Activation Mapping (Grad-CAM), a finding we support with side-by-side qualitative comparison. CerebAI additionally incorporates a native DICOM processing pipeline to facilitate future clinical translation. Code and model weights are publicly available to support reproducibility and further research.
Chen, Z.; Luo, Q.
Show abstract
Protein function prediction traditionally relies on structured gene ontology (GO) labels or multi-label classifiers. However, these labels or classifiers cannot flexibly describe molecular function, biological process, cellular component, and free-text functional narratives in a single output. In comparison, generation-based approaches offer an intuitive paradigm for flexible free-text protein annotation, with large language models (LLMs) as a representative method for protein-text modeling. Recent efforts on utilizing LLMs for protein semantic understanding and annotation generation have adopted sequence-only encoding or sequence-text contrastive alignment paradigms, yet without explicit consideration of three-dimensional structural information. To address these limitations in current protein function prediction methods, we present ProtBLIP2-SST, a two-stage framework built on the BLIP2 model architecture that bridges protein sequence, structure, and text for open-ended protein functional caption generation. Specifically, we first integrate sequence and structure information through SaProt, a protein language model (PLM) with a structure-aware vocabulary that fuses residue tokens with Foldseek-derived 3Di structural tokens. To empower the LLM to understand protein semantics, we employ a Q-Former (a querying transformer in BLIP2) with learnable query tokens as the cross-modal projector to align protein features from the frozen SaProt encoder and text features from a frozen BiomedBERT via protein-text contrasting, protein-text matching, and protein captioning objectives. After alignment, the protein features are linearly projected and prepended to the prompt embeddings of the LLM for protein captioning fine-tuning with LoRA. Trained on 441k protein-text pairs from Swiss-Prot with corresponding structures from the AlphaFold Database, our ProtBLIP2-SST outperforms sequence-only and sequence-text alignment baselines on protein captioning metrics, with ablation studies demonstrating the effectiveness of integrating structure with sequence information for improved protein understanding. Through a unified two-stage alignment-and-generation pipeline, ProtBLIP2-SST integrates protein sequence and structural information, overcomes the rigidity of traditional GO-centric classification, generating open-ended captions that jointly describe molecular function, subcellular location, and homology context in one single output.
Liu, T.; Wu, Y.; Bao, Y.; Li, W.; Li, C.; Liu, Z.; Lin, G. N.
Show abstract
Precise early diagnosis and progression prediction of Alzheimer's Disease (AD) are critical for optimizing clinical intervention. However, current methodologies often suffer from the passive utilization of clinical priors and rigid modal fusion strategies, failing to capture the heterogeneous variations of imaging biomarkers. Furthermore, predicting the precise time-to-conversion from Mild Cognitive Impairment (MCI) to AD remains a formidable challenge. To address these limitations, we propose COMPASS, a clinical-guided multi-modal framework that unifies diagnosis with a comprehensive survival strategy. Specifically, we instigate a paradigm shift to "clinical-prior-driven" learning by incorporating Clinical-Guided Spatial Attention (CGSA), which actively transforms clinical states into visual signals to modulate neural focus on pathological regions. To bridge the semantic gap between modalities, we introduce Reciprocal Semantic Interaction (RSI) via cross-attention, while a Disease-Stage-Aware Modal Fusion (DSAMF) module dynamically adjusts modal weights based on inferred disease severity to mimic clinical reasoning. Moreover, we specifically design a Dual-Head Joint Survival Risk and Time Prediction Network (DH-Net) to jointly perform quantitative conversion time prediction and patient risk stratification. Extensive experiments demonstrate that COMPASS outperforms state-of-the-art methods, achieving 83.19% accuracy in pMCI vs. sMCI classification, an MAE of 7.96 months for conversion time prediction, and a C-index of 0.819. Furthermore, we conducted in-depth neurobiological interpretability analyses, revealing right hippocampal dominance and synergistic regional impairment patterns, thereby providing new biological insights for early AD diagnosis and subtype identification.
Jiang, H.; Yang, C.; Qin, M.; You, J.; Feng, J.; Yu, J.-T.; Cheng, W.; Gong, W.
Show abstract
Precision medicine faces a critical challenge in translating high-dimensional omics data into robust disease predictions across diverse populations. Current approaches often fail under distribution shifts, partly due to their inability to encode complex biological feature dependencies. We present OmicFormer, a Transformer-based architecture that embeds two complementary statistical priors, i.e., feature-label associations and feature-feature dependencies, directly into its representation learning. This design captures local and long-range omic interactions often missed by conventional methods. Analyzing 500,000 UK Biobank participants, OmicFormer significantly outperforms strong baselines across 450 disease and 900 trait prediction tasks , with substantial gains spanning diverse metabolic, neurological, cardiovascular, and gastrointestinal conditions, alongside enhanced prediction of circulating metabolites, bone density traits, and retinal imaging biomarkers. Crucially, OmicFormer demonstrates robust generalization, achieving a substantial improvement over tree-based methods in an independent proteomics cohort across 19 diseases (GNPC, N=7,289), and outperforming tree-based models across 50 multi-site neuroimaging sites (N=4,728) for autism and schizophrenia classification. By explicitly embedding statistical structure, OmicFormer provides an interpretable and generalizable foundation for omics-based precision medicine.
Bit, S.; Guney, O. B.; Jia, S.; Kolachalama, V. B.
Show abstract
Automated interpretation of neuroimaging studies requires simultaneous assessment of multiple imaging evidence variables, each tied to distinct anatomical structures. Vision-language models (VLMs) offer a unified framework for multi-task analysis, but adapting pre-trained VLMs remains challenging. Full fine-tuning is computationally prohibitive, and joint multi-task training requires simultaneous access to all task data, which is often infeasible in clinical settings. Although model merging enables multi-task composition without joint re-training, existing methods focus on post-hoc algorithms with limited extension to VLMs and minimal application to neuroimaging. Here, we present GRadient-guided Adapter Merging (GRAM), a layer-selective low-rank adaptation (LoRA)-based fine-tuning and merging framework for multi-task neuroimaging visual question-answering (VQA). GRAM uses a gradient ratio that contrasts class-specific gradients to identify task-discriminative layers, and applies subspace-constrained projected gradient descent to restrict LoRA updates to directions consistent with the geometry of the pre-trained model. We leveraged a structured VQA benchmark, developed from the National Alzheimer's Coordinating Center (NACC) dataset, that pairs multi-sequence brain MRI studies with question-answer pairs across clinically relevant imaging evidence variables. Experiments on the VQA benchmark showed that GRAM outperformed or matched all-layer LoRA fine-tuning and a standard merging baseline while reducing inter-task interference during merging, and approached or surpassed the performance of joint multi-task training without joint re-training.
Xiao, C.; Ding, Y.; Bian, H.; Chen, Y.; Wei, L.; Zhang, X.
Show abstract
Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we propose CellTok, a tokenized single-cell language modeling approach that converts transcriptomic profiles into compact cellular token sequences and incorporates them into the vocabulary of a pretrained LLM. By representing cells as native tokens, CellTok enables cellular measurements, textual instructions, biological context, and multi-cell populations to be jointly processed within the same autoregressive modeling framework. Across diverse tasks, CellTok enable LLMs to recognize individual cells, interpret homogeneous and heterogeneous cell populations, infer disease-associated cellular states, predict cell-cell communication, model developmental trajectories, and generate cellular states. Moreover, prompt-based experiments show that providing appropriate biological context improves performance, indicating that CellTok can leverage LLM knowledge and contextual reasoning to support cellular data interpretation. These results demonstrate that single-cell transcriptomes can be transformed from a foreign molecular modality into a native language for LLMs, establishing a unified interface for modeling cells, populations, and biological knowledge in a shared token space.
Fan, Y.; Liu, X.; Wang, Y.; Zeng, Z.; Li, L.; Qiu, X.; Li, Y.
Show abstract
Embryonic development is orchestrated by complex gene regulatory networks, and learning regulatory dynamics from developmental data could allow us to understand, predict, and ultimately engineer cell fates. Here we introduce Navigo (https://github.com/aristoteleo/Navigo-release), a biologically grounded generative modeling framework that learns a developmental vector field by integrating flow matching at the population level with RNA kinetics modeling at the molecular level. Navigo accurately maps developmental trajectories across lineages on a mouse embryogenesis scRNA-seq atlas spanning 43 time points and comprising 12.4 million cells. Applied to cardiac development, Navigo enables disease modeling by mechanistically resolving regulatory networks that distinguish congenital heart disease subtypes. Navigo also predicts perturbation effects in a zero-shot manner, as validated on independent in vivo data from six knockout genotypes without perturbation-specific training, uncovering lineage-specific gene-compensation mechanisms. Moreover, Navigo guides rational cell-fate engineering, exemplified by fibroblast reprogramming analyses, including identifying pro-fibrotic barriers to cardiac fates and evaluating hundreds of pairwise transcription factor combinations for neuronal fate, each consisting of one bHLH factor and one POU factor. Overall, Navigo provides a generalizable AI platform for perturbation-effect prediction, disease modeling, and rational cell-fate engineering, advancing toward AI-based virtual embryos for developmental biology and regenerative medicine.
Pande, A.; Uyar, B.; Akalin, A.
Show abstract
Public bulk RNA-seq repositories contain hundreds of thousands of samples, creating opportunities for large-scale representation learning, but integration across studies remains challenging because of heterogeneous annotations, experimental protocols, and technical variation. While pre-trained foundation models are now widely available for single-cell RNA-seq, comparable resources for bulk RNA-seq remain scarce, motivating a model that learns a unified, tissue-aware representation directly from bulk data. We trained a supervised variational autoencoder (VAE) on a compendium of 118,263 bulk RNA-seq samples that we assembled from TCGA, GTEx, and ARCHS4 and mapped to 42 tissue categories. The model classifies tissue of origin at 94.9% balanced accuracy (weighted F1 96.2%) and compresses 16,115 genes into a 121-dimensional latent space. Tissue identity is the primary organizing axis of the latent space, while source effects remain secondary. To assess the impact of data volume, we constructed training sets at three different scales (38K, 75K, and 118K samples). Our results demonstrated that reconstruction fidelity improved incrementally with each expansion of the dataset, but with diminishing returns. We validated the model on an independent cohort of 734 paediatric tumour samples from TARGET, achieving 84.6% agreement with the expected tissue of origin. The trained model and code are available at GitHub (https://github.com/BIMSBbioinfo/flexynesis_tissue_vae_manuscript) with an interactive web application.
Wei, Y.; Eberini, I.; Meyer, F.
Show abstract
Protein thermostability is a critical property for both industrial and biomedical enzyme applications, yet experimental evaluation of mutation-induced stability changes remains laborious and costly. Here, we present ThermoFusion, a hybrid deep learning framework that integrates 3D protein structure embeddings from ThermoMPNN with sequence-based embeddings from the pretrained protein language model ESM2 to predict the effects of single-point mutations on protein stability ({Delta}{Delta}G). ThermoFusion exhibits robust generalization, maintaining high predictive accuracy across out of distribution sequences with low identity to the training set -- a scenario where many other machine learning models, including ThermoMPNN and state-of-the-art tools, perform poorly due to reliance on memorization. Benchmarking on a curated enzyme dataset comprising of 105 enzymes and 3144 mutations shows that ThermoFusion reliably identifies stabilizing mutations while accurately predicting stability for enzymes beyond its training set. These results establish ThermoFusion as a powerful tool for rational enzyme design beyond its training set.
Pareja-Lorente, E.; Aloy, P.
Show abstract
Foundation models have emerged as powerful tools for learning transferable representations of biological systems, yet their latent spaces are typically optimized to capture cellular state rather than the effects of perturbations. Here, we demonstrate that a biological foundation model can be repurposed to learn a fundamentally different representation by changing its learning objective. We finetuned scGPT, a transformer pre-trained on over 30 million single cell transcriptomes, on more than three million LINCS L1000 perturbation profiles using a supervised objective that predicts perturbation identity. This transformed the latent space into a perturbation centric representation that aligned transcriptional responses induced by the same chemical or genetic perturbation across heterogeneous experimental conditions. Finetuned embeddings substantially outperformed both gene expression profiles and the original pretrained model, recovering [85, 100%] of perturbations within the top 100 nearest neighbors and increasing perturbation classification accuracy from [10, 19%] to [25, 49%]. Remarkably, although the model was trained exclusively to recognize perturbation identity, the learned representation spontaneously captured orthogonal biological relationships never provided during training, including chemical similarity (AUROC up to 0.81), mechanisms of action (Hit@10 up to 100%), compound target relationships (AUROC up to 0.74), and functional relationships between genetic perturbations. The resulting embedding space enabled mechanism-of-action annotation of nearly 12,000 previously uncharacterized compounds, prioritization of target related chemical genetic associations, and contextualization of unseen perturbations and external transcriptomic datasets. Together, our results establish objective-driven adaptation as a general strategy for repurposing biological foundation models to learn reusable representations of complex biological phenomena.